BMC Medical Research Methodology
○ Springer Science and Business Media LLC
Preprints posted in the last 7 days, ranked by how well they match BMC Medical Research Methodology's content profile, based on 47 papers previously published here. The average preprint has a 0.06% match score for this journal, so anything above that is already an above-average fit.
Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.
Show abstract
We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.
Xiang, S.; He, H.; Xie, Z.; Cheng, C.-Y.; Li, H.; Liu, D.
Show abstract
Agentic workflows can coordinate modelling, but balancing predictive performance, measurement burden and reproducibility is unclear. We developed DXA Agent, an agentic workflow for dual-energy X-ray absorptiometry (DXA) outcomes integrating planning, feature-model refinement, tools, provenance and hypothesis-generating interpretation. Models were independently developed and tested in UK Biobank (5,318 participants) and the National Health and Nutrition Examination Survey (NHANES; 3,777 participants), using cost-efficient and no-limit strategies. Across 20 UK Biobank and three NHANES bone mineral density sites, cost-efficient models achieved lower RMSE and higher R2 than the best conventional comparator, with median relative RMSE reductions of 10.9% and 9.9%, respectively. Classification was task dependent: UK Biobank osteoporosis averaged AUROC 0.839 and PR-AUC 0.182, whereas NHANES performance was comparable with conventional models. Higher-burden features did not consistently improve prediction. These retrospective, cohort-internal findings position DXA Agent as an inspectable, measurement-burden-aware research workflow requiring independent prospective validation.
Mamiya, H.; Zhang, Q.; Zhang, X.; Yan, Y.; Sharma, A.
Show abstract
Wearable (accelerometer) data and machine-learning allow objective assessment of the amount of daily physical activity. However, wearable-derived human activity is subject to measurement error. No studies have corrected the dose-response association between physical activity and survival time to chronic diseases, including cardiovascular disease (CVD). The objective is to estimate the measurement error-corrected association between CVD events and multiple measures of daily duration of light and total physical activity, derived from machine-learning and conventional accelerometer-processing methods. Our method combined an accelerated failure time model, spline, and simulation-extrapolation (SIMEX). The method recovered the true dose-response non-linear association in simulated data, while the naive model failed to capture it due to substantial attenuation. Application to the UK Biobank accelerometer cohort also showed an increased protective association of total physical activity after SIMEX correction (Time Ratio [TR] = 1.56, 95% CI: 1.28-1.82 vs. TR = 1.38, 95% CI: 1.24-1.54 for SIMEX-corrected vs. uncorrected dose-response association between the 95th and 5th percentiles of total activity), with a similar increase for light physical activity. Sensitivity analysis indicates that the female population experiences a substantially larger protective association after SIMEX correction than males. Dose-response survival analysis is a widely used analytical method in physical activity epidemiology and benefits from measurement error correction.
Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.
Show abstract
Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.
De Luca, S.; Fava, C.; Rizzo, G.; Visconti, A.; Berchialla, P.
Show abstract
Background. Patient stratification from multi-omics and clinical data is essential for uncovering disease heterogeneity and moving toward more personalized treatment strategies. However, integrating heterogeneous data layers while identifying robust patient strata remains challenging. Methods. We introduce Reduced Fusion of Multi-Omics Stratification (RedFuMOS), a novel three-step approach for patient stratification based on mixed-type multi-omics data. RedFuMOS extends Similarity Network Fusion to accommodate mixed-type data layers and layer-specific similarity measures for data integration, includes a dimensionality reduction step to mitigate the curse of dimensionality, and performs patient stratification using density-based hierarchical clustering with HDBSCAN. It also implemented an automated optimization procedure to identify the best set of hyperparameters, minimizing the need for manual tuning. Results. RedFuMOS outperformed six state-of-the-art tools for multi-omics patient stratification in a comprehensive simulated benchmarking study, which also confirmed that, although computationally expensive, the dimensionality reduction step is crucial for achieving good stratification performance. Additionally, RedFuMOS identified two clinically relevant patient strata in a small real-world cohort of patients with Philadelphia chromosome-positive chronic myeloid leukaemia. Conclusion. RedFuMOS provides a flexible framework for integrating heterogeneous multi-omics and clinical data. RedFuMOS is available as an R package at http://github.com/delucasara/RedFuMOS.
Yano, Y.; Nagasu, H.; Hiroshi, K.; Ohashi, M.; Isaka, Y.; Okada, H.; Nangaku, M.; Kashihara, N.
Show abstract
Background: Traditional real-world studies comparing SGLT2 and DPP4 inhibitors on renal outcomes rely on propensity score matching, which causes high-dimensional data loss. We used causal machine learning (Causal ML) to unmask heterogeneous treatment effects in diabetic kidney disease (DKD). Methods: Using data from 4,588 patients within the Japanese J-CKD-DB-Ex registry, we implemented a doubly robust (DR) learning framework (Linear DR-learner with XGBoost) to compare SGLT2 and DPP4 inhibitors. Outcomes included the chronic eGFR slope and a composite renal endpoint ([≥] 50% eGFR decline or end-stage kidney disease). Heterogeneity was explored via causal SHAP and decision trees. Results: At the population level, SGLT2 inhibitors modestly slowed chronic eGFR decline (average treatment effect [ATE] = 0.14 [95% CI: -0.86, 1.15] mL/min/1.73m^2/year) and reduced composite endpoint risk by 9% (ATE: -0.09 [-0.11, -0.08]) versus DPP4 inhibitors. However, individual-level counterfactual analysis suggested that for the chronic eGFR slope, non-glinide users with stable pre-treatment trajectories who were also taking ACE inhibitors had a greater benefit from SGLT2 inhibitors (ATE: 2.95 [-0.68, 6.58]). Conversely, glinide users with steep pre-treatment decline had a greater benefit from DPP4 inhibitors (ATE: -8.98 [-16.11, -1.85]). For composite renal events, SGLT2 inhibitors had a 28% absolute risk reduction within the algorithmically identified high-risk subgroup (eGFR [≤] 28.1 mL/min/1.73 m^2 and positive proteinuria; ATE: -0.28 [-0.33, -0.23]). Even non-proteinuric decliners demonstrated a 8% risk reduction with SGLT2 inhibitors (ATE: -0.08 [-0.10, -0.06]). Conclusion: Causal ML advances precision medicine in DKD, shifting from uniform prescribing to individualized, data-driven therapy targeting distinct intrarenal pathways.
Jaber, A.; Hughes, L.; Cameron, A. C.; Quinn, T. J.
Show abstract
Background: Systematic reviews of clinical prediction models increasingly include studies using artificial intelligence (AI) and machine learning (ML) methods alongside traditional multivariable regression approaches. A previously published Excel tool enabled standardised data extraction using the CHARMS checklist and risk of bias assessment using PROBAST. The recent publication of the PROBAST+AI framework, which distinguishes the assessment of model development quality from the assessment of model evaluation risk of bias and assesses applicability in both parts, necessitates an updated digital instrument applicable across prediction modelling methods. Methods: We updated an open-access Excel tool to incorporate the full PROBAST+AI framework. The updated template incorporates structural separation between assessment of model development quality and model evaluation risk of bias, with applicability assessed in both parts. It also incorporates updated signalling questions, including those addressing methodological issues particularly relevant to AI/ML, and automates the generation of summary tables and graphical displays. Results: The updated tool (CHARMS & PROBAST+AI Template) contains 11 worksheets and supports data extraction and appraisal for up to 30 prediction models. Dedicated, linked worksheets enable separate assessment of model development and model evaluation, with Domain 4 distinguishing among Apparent, Internal, and External evaluation settings. Key updates include dedicated assessments for predictor pre-processing, class imbalance handling and recalibration, data leakage prevention, and replication of the full model development pipeline within resampling procedures. Automated sheets dynamically format tables and summary charts covering PROBAST+AI parts. Conclusions: The CHARMS & PROBAST+AI Excel template provides a standardised, user-friendly, and rigorous digital framework for systematic reviewers appraising traditional statistical and AI-driven clinical prediction models.
Luna-Martinez, N.; Cruz-Rodriguez, E. X.; Bernal-Castro, E. A.
Show abstract
Background Dengue is a major public health challenge, and predictive models are crucial for early warning systems. However, many current modeling practices rely exclusively on climatic factors or employ complex algorithms that lack the interpretability needed for informed public health decision-making. To address these shortcomings, we developed and validated a multidimensional, interpretable statistical model to predict monthly dengue incidence. Methodology/Principal Findings We used a Generalized Linear Mixed Model (GLMM) with a Negative Binomial distribution to analyze 14 years (2010-2023) of spatiotemporal data from 37 municipalities in Huila, Colombia, an endemic region. The model integrates non-linear and lagged effects of climatic, demographic, and socioeconomic factors. The final model underwent rigorous external validation on an independent test set (2021-2023). Our model demonstrated high predictive discrimination (R2 = 0.743, Spearman's {rho} = 0.657), accurately capturing the timing of epidemic outbreaks. Key findings include the identification of an optimal thermal window for transmission at 27-28{degrees}C, a threshold effect for precipitation above 800 mm, and a saturation dynamic in outbreak autocorrelation. Conclusions/Significance This mechanistically-informed statistical approach provides a robust and transparent tool for epidemiological surveillance, successfully balancing high predictive performance with the explanatory power needed for effective, data-driven public health interventions.
Okundaye, D. O.; Isiekwene, C. C.
Show abstract
Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.
Mengi, A.; Bagita-Vangana, M.; Tesine, P.; Laman, M.; Bolnga, J. W.; Ome-Kaius, M.; Kulimbao, J.; Mase, J.; Mal, L. S.; Mnjala, H.; Lee, G.; Cassidy-Seyoum, S. A.; Thriemer, K.; Unger, H. W.
Show abstract
Disseminating study results to participants is an ethical responsibility for researchers but remains uncommon in low- and middle-income countries, and participants preferences for receiving study results are poorly understood. This study examined study result dissemination preferences among pregnant women in a phase III malaria prevention trial in Papua New Guinea (PNG). Participants completed an interviewer-administered questionnaire (survey) assessing their interest in and motivation for receiving trial results and preferences for dissemination methods and content. Associations between participants characteristics and dissemination preferences were explored using multivariable logistic regression analysis. Of 1172 trial participants, 96.0% (1125/1172) completed the survey, and of these 99.6% (1121/1125) wanted to learn about the trial results. The main motivation factors driving participants interest were an acknowledgment of their contribution to research (51.7%; n=579) and a better understanding of the study (45.0%; n=505). Most participants (78.9%; n=884) wanted to learn about the trial findings through written summary and a group meeting with other participants at the nearest clinic (31.1%, n=349). Multivariable regression analysis indicated that participants from rural/peri-urban clinics were more likely to choose non-electronic media dissemination approaches such as a group meeting as compared to urban-dwelling participants. Frequently selected items (>50% of participants) for content included information regarding good results of the study, purpose of the study, medical treatment advances, results specific to me, and how study was conducted. There was heterogenicity in the preference for dissemination content: compared to urban clinics rural clinics are less likely to want to learn about how and why study was conducted and medical and scientific advances. Overall, the majority wanted to learn about trial results, highlighting the importance of integrating dissemination into research activities in PNG. Variation in preferences for mode and content of dissemination between study clinics suggests that dissemination activities could be tailored to local context and preferences.
de Araujo Morais, J. H.; Dias Ferreira, C.; Saraceni, V.; Medeiros de Oliveira Cruz, D.; Mateus Oliveira Aguilar, G.; Cruz, O. G.
Show abstract
Motivation: With the scaling frequency and intensity of extreme heat events across the globe, it is critical for public institutions to develop early detection systems and continuous monitoring of these events and their impacts. In Brazil, Rio de Janeiro was the first city to publish its heat protocol, with the Rio Heat Dashboard as a central component of this system. Implementation: The dashboard was implemented using R/Shiny and integrates climatic and health data from multiple sources. General features: The application comprises real-time heat exposure monitoring and automatic alert level classification, which is monitored daily by multiple municipal actors and supports activation of actions specified in the heat protocol. It also features a health impact module, which lists each heat event and its impact on mortality, and primary care and emergency visits. Availability: The source for full reproducibility is available through https://github.com/joaohmorais/RioHeatDashboard.
Li, S.; Zhang, W.; Xing, X.; Shen, Z.; Wang, Y.; Chen, Z.; Neto, O.; Yu, Y.; Wu, C.; Lin, L.
Show abstract
Background Late-stage cancer incidence is being considered as an earlier endpoint in cancer-screening trials, but its trial-level association with cancer-specific mortality may depend on evidence selection and endpoint harmonization. We evaluated the robustness of this association to source-verified additions. Methods We reconstructed the PubMed corpus underlying a 41-comparison review. Gemini 3.1 Pro Preview was used only to prioritize reports for blinded human reassessment. Reviewers determined eligibility, linked reports from the same trial, harmonized endpoints, and verified comparison-level data. We recalculated unweighted Pearson correlations overall and by cancer type after adding earliest-compatible trial comparisons. Results Among 1209 candidate records, 996 PDFs were assessed. Thirty-three reports absent from the source review were prioritized; 26 were eligible, representing 18 trials, and 8 provided compatible comparisons. Adding these comparisons increased the dataset from 41 to 49 and attenuated the overall correlation from 0.73 (95% confidence interval [CI] = 0.55 to 0.85) to 0.59 (95% CI = 0.37 to 0.75). Updated correlations were 0.49 (95% CI = -0.26 to 0.87) for breast, -0.23 (95% CI = -0.71 to 0.40) for colorectal, and 0.83 (95% CI = 0.54 to 0.95) for lung cancer. One sparse-event comparison influenced the colorectal estimate. Conclusions The overall association was sensitive to evidence composition, and cancer-specific stability varied. Late-stage incidence should be evaluated by cancer type and with prespecified sensitivity analyses for evidence selection and endpoint definitions. Model-assisted prioritization cannot replace human eligibility review, trial reconciliation, and source verification.
Yakubu, S.; Mousavi, S.; Eden, J.; Kabajulizi, J.; Palade, V.; Daneshkhah, A.
Show abstract
Communities exposed to flooding can experience markedly different mental health outcomes, yet conventional resilience indicators capture only part of the social and contextual conditions that may explain this variation. This study develops a multilevel and predictive framework for examining community resilience and depressive symptoms following flood exposure in Indonesia. Data were drawn from 20,303 respondents aged 15 years and older nested within 312 communities in the Indonesia Family Life Survey (IFLS-5). Depressive symptoms were assessed using the 10-item Centre for Epidemiologic Studies Depression Scale (CES-D-10), with Rasch Partial Credit Model calibration used to examine measurement properties. Bayesian multilevel models quantified between-community heterogeneity and assessed how far observable structural resources accounted for this variation. Community resilience was represented through two complementary constructs: structural resilience, based on observable socioeconomic and social-capital resources, and Latent Community Protective Capacity (LCPC), a model-derived proxy for residual contextual variation in depressive-symptom risk. Approximately 6 percent of variation was attributable to between-community differences, while observable structural resources explained only part of this heterogeneity. Structural resilience and LCPC were weakly correlated (r = 0.155). Moderation analyses provided no clear evidence that structural resilience altered the flood-depression association, while LCPC showed a directionally consistent but uncertain buffering pattern. Predictive models incorporating community-level information improved discrimination, with the best-performing model reaching an ROC-AUC of approximately 0.71. The findings suggest that observable resource-based indices provide an incomplete account of community-level mental health vulnerability and that residual contextual measures may provide complementary information, while requiring cautious interpretation and independent validation.
Chin, A. T.; Zhu, N.; Vangala, S.; Woo, H.; Wisk, L. E.; Kingsley, T.; Mafi, J. N.; Lukac, P. J.
Show abstract
BACKGROUND Generative AI (genAI) chart summarization tools embedded in electronic health records (EHRs) are being rapidly deployed across U.S. health systems. Although these tools represent a promising solution to alleviate cognitive burdens, their effects have not been examined in randomized-clinical trials (RCTs). METHODS In this pragmatic RCT at a single academic health system, 284 outpatient clinicians across forty-two specialties were assigned 1:1 to Epic's outpatient chart summarization tool or a usual-care control arm over 90 days, from February 23 to May 23, 2026. The primary outcome was physician task load (PTL) adapted for pre-charting. Prespecified exploratory outcomes included additional validated psychometrics as well as usability, safety, and time-based measures. Descriptive statistics included interaction and usage of the tool. RESULTS Of 74,474 AI chart summaries generated, 14.2% were interacted with by a clinician; the proportion of generated summaries interacted with declined from 21.5% in month 1 to 10.5% in month 3, and the proportion of clinicians using the tool at least once per month declined from 88.7% to 66.2%. The adjusted between-arm difference in PTL at follow-up favored the intervention arm (scale 0-400; -27.4; 95% CI, -49.4 to -5.3; P=0.02). Among the Professional Fulfillment Index (PFI; scale 0-4, lower=better) psychometrics, overall burnout (-0.20; 95% CI, -0.38 to -0.01) and work exhaustion (-0.24; 95% CI, -0.47 to -0.02) were lower in the intervention arm, with little difference in overall professional fulfillment (+0.04; 95% CI, -0.16 to 0.25). Charting time per encounter showed no significant between-arm difference during steady state (-1.2 seconds; 95% CI, -19.0 to 16.6). The net promoter score was -22, indicating that on average, clinicians did not recommend the tool. Among free-text respondents, 57.1% reported at least one concern, most commonly tool limitations or inaccurate information. No adverse patient safety events or near-misses were reported. CONCLUSION An EHR-integrated AI chart summarization tool modestly reduced physician task load and was associated with lower burnout, without time savings and against declining engagement. Sustained usage and oversight of reported inaccuracies remain open challenges.
ye, y.; Zeng, Z.; Tian, X.; Yuan, Z.; Wang, J.; Zhu, Y.
Show abstract
Artificial intelligence applied to routine electrocardiograms (ECGs) has largely focused on detecting existing disease or predicting individual cardiovascular outcomes. Whether ECGs can support prediction of multiple future diseases across organ systems remains unclear. We developed ECG-RISK, a multitask survival model for 67 incident three-character ICD-10 endpoints using ECG waveforms, demographic characteristics and routinely collected laboratory data from 86,673 MIMIC-IV patients. Discrimination was highest for heart, brain, kidney and lung endpoints, with organ-level C-indices ranging from 0.796 to 0.825, whereas liver and pancreatic endpoints showed lower discrimination. The ECG-only model achieved strong discrimination across most endpoints, whereas the incremental improvement gained by incorporating ECG and laboratory inputs beyond demographic information varied substantially across endpoints. Across the nine exploratory aggregated outcomes, Kaplan Meier curves showed clear separation among model-score tertiles. Discrimination was highest for dementia (C-index, 0.891) and heart failure (C-index, 0.857). These findings support the feasibility of ECG-based longitudinal risk prediction across multiple diseases. External validation and competing-risk analyses are required to assess generalisability and clinical utility.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
Chu, W.-Y.; Alves, F.; Dorlo, T. P. C.
Show abstract
Introduction Miltefosine is the only approved oral antileishmanial agent, but its use in women of childbearing potential (WOCBP) is restricted due to preclinical teratogenicity. Current labeling recommends contraception during treatment and for at least five months thereafter, based solely on its long terminal elimination half-life. This study re-evaluated the required contraceptive duration using an exposure margin-based approach. Methods Virtual populations were generated from anthropometric data of 382 Indian, 4,462 Eastern African, and 4,019 Brazilian WOCBP with leishmaniasis. Published population pharmacokinetic models were used to simulate miltefosine exposure following 14-42-day regimens for visceral leishmaniasis (VL), post-kala-azar dermal leishmaniasis (PKDL), and cutaneous leishmaniasis (CL). A developmental safety exposure threshold was derived from the rat no-observed-adverse-effect level (0.6 mg/kg/day for 10 days) and adjusted using a 10-fold safety margin. Contraceptive durations resulting in median residual post-contraception exposure (AUCEOC-{infty}) below this threshold were considered supportive of contraceptive discontinuation. Results The developmental safety exposure threshold was estimated at 2.5 mg{middle dot}day/L. Despite pharmacokinetic differences across geographical regions and disease manifestations, required minimum contraceptive durations were consistent: three months from treatment initiation for the 14-day regimen, four months for 21- and 28-day regimens, and five months for the 42-day regimen. For the 14-day regimen, a single dose of a long-acting injectable contraceptive administered at treatment initiation would provide sufficient coverage. Conclusion An exposure margin-based approach supports shorter contraceptive durations than current recommendations. For the 14-day VL regimen, three months of contraception may provide a practical alternative to current labeling.
Sadia, H.; Doyon, N.; Duchesne, S.
Show abstract
Background Understanding the mechanisms underlying brain aging and age-related pathological changes is essential for advancing brain health research. Our group previously developed a mechanistic mathematical model of healthy brain, Chamberland et al. (2024) that integrates key biological processes involved in normal aging, from which Alzheimer's disease (AD) related changes may emerge naturally. Objectives To characterize and validate this brain model by evaluating its sensitivity, calibrating its parameters, and assessing generalizability in independent populations. Methods The model represents the evolution of key biological processes associated with brain aging, including amyloid beta (A{beta}), tau pathologies, neuroinflammation, and neuronal death. After identifying the 30 most influential parameters, we calibrated the model using cognitively normal (CN) participants from the AD Neuroimaging Initiative (ADNI) database (n = 211) by minimizing a loss function composed of three outcomes (AB) plaques, tau tangles, and neuronal density). The calibrated model was then applied to the UK Biobank cohort (n = 35,899) of normal controls (aged 44-82 years). The effects of sex and APOE were evaluated using stratified simulations. Results Parameter calibration significantly reduced the prediction errors for A{beta} and tau. Neuronal density predictions showed strong agreement in the UK Biobank cohort. The variance decomposition identified APOE status as a major contributor to variability in A{beta}. Conclusion Our validated brain health model links mechanistic pathways with population data and reproduces neuronal density patterns in an independent cohort. These findings support its use as a framework for studying brain aging and investigating how Alzheimer's disease related pathological changes may emerge with aging.
Marban-Castro, E.; Muhwava, L.; Girdwood, S.; Kemp, T.; Freitas, J.; Kamau, Y.; Otieno, M.; Akach, D.; Morato, A.; Sanz, S.; Fiechter, V.; Erkosar, B.; Watson, M.; Vetter, B.; Haldane, C.; Shilton, S.; Rheeder, P.; Dave, J. A.; Carrihill, M.; Karsas, M.
Show abstract
Introduction: Continuous glucose monitoring (CGM) offers an advancement over traditional self-monitoring of blood glucose (SMBG) for people living with type 1 diabetes (T1D). However, evidence on the acceptability and feasibility of different CGM use cases in African populations remains limited. Methods: This was a pragmatic three-arm, randomised controlled trial on CGM conducted among people living with T1D in three public healthcare clinics in South Africa. Participants were assigned to Arm 1 (continuous CGM), Arm 2 (periodic CGM), or Arm 3 (SMBG). Diabetes education was provided at all study visits. Feasibility was assessed by adherence to CGM use and through the Glucose Monitoring Satisfaction Survey (GMSS). Diabetes distress was measured by the Diabetes Distress Scale (DDS), health-related quality of life (HRQoL) by the EQ-5D scales, and acceptability using the Theoretical Framework of Acceptability (TFA). Surveys were collected on paper and transferred to OpenClinica. Analyses were performed in R. The trial was registered in the Clinical Trials Registry (NCT05944718) on July 13, 2023. Results: A total of 83 participants were included in Arm 1, 85 in Arm 2, and 80 in Arm 3. CGM mean active time was 55% in Arm 1 versus 69% in Arm 2. The proportion of participants meeting the [≥]70% active time threshold was higher in Arm 2 (52%) than in Arm 1 (34%). Diabetes' distress declined across arms during the intervention period, with no significant difference between arms; distress increased slightly six months post-intervention but remained below baseline. At 6 months, glucose monitoring satisfaction was significantly higher in both CGM arms than in the SMBG arm, and satisfaction increased over time in CGM arms. Health-related quality of life remained stable across arms during the intervention period with no significant difference between arms. High acceptability was observed in both CGM arms, with higher ratings in the periodic arm. Conclusions: CGM was acceptable to people living with type 1 diabetes and feasible to use in public-sector clinics in South Africa, with high acceptability under continuous and periodic use. Health-related quality of life remained stable across arms, and diabetes-related distress declined, during the intervention period, across arms. Glucose monitoring satisfaction rose significantly in both CGM arms compared to SMBG. Periodic CGM might be a promising and potentially more scalable option than continuous use for public-sector care.